Papers with qualitative error analysis

9 papers
A Study on Entity Resolution for Email Conversations (2020.lrec-1)

Copied to clipboard

Challenge: This paper addresses the task of entity resolution in email conversations.
Approach: They propose to create an annotated seed corpus of email threads labeled with entity coreference chains and evaluate their models for the task.
Outcome: The proposed model performs well on the entity resolution task for email conversations.
Medical Spoken Named Entity Recognition (2025.naacl-industry)

Copied to clipboard

Challenge: Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc.
Approach: They present a spoken NER dataset in the medical domain using pre-trained models that are encoder-only and sequence-to-sequence.
Outcome: The dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.
A Multi-Orthography Parallel Corpus of Yiddish Nouns (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora of Yiddish text are limited to a single, potentially non-standard orthography . non-phonetically spelled Hebrew words are the largest cause of error, according to our study .
Approach: They propose a multi-orthography parallel Yiddish corpus based on Wiktionary scraping . they also demonstrate how the system can be used to bootstrap a transliteration model .
Outcome: The proposed system achieves error rates between 16.79% and 28.47% on the test set.
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets (2024.findings-eacl)

Copied to clipboard

Challenge: Using large language models (LMs) for query or document expansion can improve generalization in information retrieval.
Approach: They conduct the first comprehensive analysis of large language models (LMs) for query or document expansion.
Outcome: The proposed expansions improve retrieval performance for weaker models but harm stronger models.
CrossAligner & Co: Zero-Shot Transfer Methods for Task-Oriented Cross-lingual Natural Language Understanding (2022.findings-acl)

Copied to clipboard

Challenge: Task-oriented personal assistants enable people to interact with devices and services using natural language.
Approach: They propose a method to acquire task knowledge in a high-resource language and then transfer it to the low-resourced language(s) they use unlabelled parallel data to perform a quantitative analysis of the methods.
Outcome: The proposed methods exceed state-of-the-art (SOTA) scores across nine languages, fifteen test sets and three benchmark multilingual datasets.
BLM-s/lE: A structured dataset of English spray-load verb alternations for testing generalization in LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Current NLP models are achieving performance comparable to human capabilities on well-established benchmarks.
Approach: They propose a BLM task to identify a missing element in a linguistic pattern from a list of candidate options based on a given matrix.
Outcome: The proposed framework is based on the spray-load verb alternations in English as a case study.
Detecting Negation Cues and Scopes in Spanish (2020.lrec-1)

Copied to clipboard

Challenge: Negation is a phenomenon that "relates an expression e to another expression with a meaning that is in some way opposed to the meaning of e" previous work on negation in English has focused mostly and only recently on annotation tasks.
Approach: They propose a machine learning system that processes negation in Spanish . they use a corpus from the SFU corpus to perform two tasks .
Outcome: The proposed system outperforms state-of-the-art in negation cue detection and scope identification.
BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification (2023.emnlp-main)

Copied to clipboard

Challenge: a number of studies have tried to detect and control the spread of such abusive memes on social media platforms.
Approach: They build a Bengali meme dataset to test models for abusive memes . they find that multimodal models that use both textual and visual information outperform unimodal models .
Outcome: The proposed model outperforms unimodal models in a Bengali meme dataset.
An Experimental Study on the Influence of Culture on Cross-Lingual Sentiment Transfer (2026.acl-long)

Copied to clipboard

Challenge: Identical linguistic expressions can convey different sentiments across cultural contexts . current multilingual models often reduce language to symbolic representation . cultural misalignment is a structural bottleneck, authors say .
Approach: They conduct an empirical study to quantify the influence of culture on cross-lingual sentiment transfer across 7 common SMLMs and 5 linguistically diverse languages.
Outcome: The proposed model disentangles cultural factors from confounding variables and shows cultural distance is a negative predictor of transfer performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations